[AI-7115] Add a request retry strategy to the async GitHub client - #24963
[AI-7115] Add a request retry strategy to the async GitHub client#24963AAraKKe wants to merge 11 commits into
Conversation
- New retry.py: RetryPolicy plus composable predicates, executed by stamina. - Split the two layers: _request retries, _rate_limited_request handles rate limits. - Per-endpoint defaults by replay safety, overridable per call with retry=. - Never follow or retry an unexpected redirect; report it with the endpoint. - Retry the artifact redirect and signed download as a pair. - Expose the limits through [dispatcher.github_retries].
evalya-impact-summaryevalya impact analysis |
|
✅ All CI checks and tests passed. 🎉 All green!🧪 All tests passed 🎯 Code Coverage (details) 🔗 Commit SHA: 78bdf90 | Docs | View more details | Give us feedback! |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c86bea94d0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
- Move the retry config into dispatcher_config, next to the other config models. - Group module constants at the top of retry.py and trim the comments. - RetryPolicy is a plain class with a typed replace instead of a dataclass. - Move the client-specific guard and the retry cause into the client module. - Redact the query string from the artifact URL before logging it.
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f3b5071d14
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
httpx builds its message from the full URL, so a retryable failure from the storage host carried the presigned signature into this client's log line and stamina's retry hook. Raise without the URL instead of redacting at each sink.
There was a problem hiding this comment.
💡 Codex Review
When any HTTPX transport failure occurs, the original exception has already been associated with the outgoing request, but constructing a replacement with only a message discards that context. Callers using the standard exc.request.url pattern will therefore get RuntimeError: The .request property has not been set, and error reporting loses the request metadata; pass the original request when constructing the replacement or preserve and re-raise the existing exception.
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…logs A transport error's reason is quoted into our message and, with a chained cause, printed in full by Python. Both were URL-free only because of how httpx builds that message. Redact the query string and drop the chain instead.
Keeps the parameter names, which say which signing scheme was in play, and masks every value rather than the ones known to be secret: the parameter carrying the signature is X-Amz-Signature on S3 and sig on Azure Blob, so an allowlist would leak the first time a download redirects somewhere new.
Building a replacement exception dropped the request httpx had attached, so exc.request raised RuntimeError for a caller. Rewrite the message in place, as the artifact download already does.
|
On the Worth noting it is not introduced by this PR. The same wrapping sits at Applied the same fix to |
|
@codex review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 97f6e0ef83
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The defaults are shared for the life of the process, so tuning one in place changed every client that held it. Frozen dataclasses with __post_init__ validation are the idiomatic way to prevent that; the alternatives for a plain class cost more boilerplate for the same result.
|
|
||
| def is_redirect_status(status_code: int) -> bool: | ||
| """Whether `status_code` is a redirect, Location header or not.""" | ||
| return status_code in REDIRECT_STATUS_RANGE |
There was a problem hiding this comment.
304 status code falls under this range but it is a Not Modified, it has no Location.
We either set the status codes manually of add an exception for when it is a 304.
There was a problem hiding this comment.
Good catch. Switched the guard to httpx's has_redirect_location (301/302/303/307/308 and a Location present), so a 304 no longer reports a redirect to a Location that does not exist.
The artifact endpoint still gets any 3xx handed back, since it validates the status and the Location itself and reports a bad one more precisely.
| def with_query_masked(text: str, url: str) -> str: | ||
| """`text` with the query of `url` masked, for a message someone else built out of that URL.""" | ||
| query = url.partition("?")[2] | ||
| return text.replace(query, masked_query(query)) if query else text |
There was a problem hiding this comment.
Question: doesn't httpx re-encode the query? we match the query as a raw string under masked_query
There was a problem hiding this comment.
It does for some values: a raw space comes back as %20. The old code matched the query against the URL we passed in, so that mismatch would have silently left the signature in the message.
Dropped the matching entirely. The whole query is now replaced with *** without parsing any of it, and a status error's reason is built from status_code/reason_phrase instead of rewriting httpx's message.
| assert policy.should_retry(_status_error(502)) | ||
|
|
||
|
|
||
| def test_a_shared_default_cannot_be_retuned_in_place() -> None: |
There was a problem hiding this comment.
Suggestion: I would drop this test: it asserts dataclass(frozen=True) raises — CPython, not our code.
There was a problem hiding this comment.
Agreed, dropped. Frozen also cannot regress unnoticed: a non-frozen RetryPolicy cannot be a field default, so the module stops importing. Replaced it with a test covering replace.
…hub-client-retry # Conflicts: # ddev/src/ddev/cli/ci/tests/dispatcher_config.py
- A 304 is no longer reported as a redirect. The guard is httpx's `has_redirect_location`, so only a Location the client declines to follow raises `GitHubUnexpectedRedirectError`; the artifact endpoint still gets any 3xx back to validate the status and Location itself. - A signed URL loses its whole query instead of having values masked parameter by parameter, and a status error's reason is built from the response rather than by rewriting httpx's message. Nothing about the query is parsed, so no encoding or delimiter has to be guessed right. - Dropped the test asserting a frozen dataclass refuses assignment, which is CPython's behaviour rather than ours, and kept one covering `replace`.
| retry_policies: RetryPolicies | None = None, | ||
| logger: logging.Logger | None = None, |
There was a problem hiding this comment.
Question: These new arguments aren’t set here. I assume that’s because the PR’s aim is to update the client rather than the dispatcher, but I wanted to mention it just in case since Claude flagged it as a request.
There was a problem hiding this comment.
This is just so the client has a logger we can control, but right now everything that has to do with the logger is bound to change when we implement the monitoring part. This is all very temporary but yes, they would need to be injected.
Although I am still unsure if I will be handling it like that or through context vars... still unknown
There was a problem hiding this comment.
Suggestion: these tests exercise the retry layer they're meant to isolate from.
Before this PR, _request was the only request method — it owned rate-limit pacing and retries. This PR splits that in two: the old _request was renamed to _rate_limited_request, and a new _request was introduced above it that adds the non-rate-limit retry strategy (retry.py / stamina).
This file's job is to test the rate-limit layer in isolation, but only one test (test_the_rate_limit_layer_does_not_retry_a_transport_error, line 220) was updated to call the renamed _rate_limited_request. The other six — lines 116, 132, 163, 187, 203, and 235 — still call client._request(...), so they now run through the new outer retry guard (_refuses_retry) as well as the rate-limit logic they're supposed to be testing alone. In practice this mostly still passes today (the guard categorically refuses to retry rate-limit-confirmed responses, so the outcomes happen to line up), but it means a regression in either layer could now surface as a test failure in the other layer's file, which defeats the point of having this file separate from test_retry.py.
Suggested fix — repoint these six calls to _rate_limited_request, matching line 220:
- await client._request("GET", "/x")
+ await client._rate_limited_request("GET", "/x")at lines 116, 132, 163, 187, 203, and 235. No other changes needed — the assertions and comments at each site still describe the right behavior once they're calling the right layer.
There was a problem hiding this comment.
Thanks for the catch! renamed them so we are not using retries in there.
The file exists to test rate-limit pacing in isolation, but six calls still went through `_request`, which now adds the retry strategy on top. No behaviour changes today, since the guard refuses every rate-limit and auth failure these tests use. It matters because the file does not opt into `instant_backoff`, so anything that made the outer layer retry here would sleep on real backoff instead of failing. Also drops a duplicated assertion in the policy-tuning test.
Validation ReportAll 21 validations passed. Show details
|
What does this PR do?
Adds a retry strategy to the async GitHub client for the failures that are not rate limiting, and separates it from the rate-limit handling that was already there.
Structure worth knowing before reading the diff:
_request(retries) wraps_rate_limited_request(today's loop, renamed). That order matters: each retry re-acquires the limiter, so it waits out any pause the governor is holding. Rate-limit responses stay owned by the inner layer and are never retried by the outer one.retry.pydescribes, stamina executes.RetryPolicyis data: what to retry on, how many attempts, what backoff. No sleeping or backoff arithmetic of ours.retry=to override, and policies compose.download_artifactretries as a pair. The signed URL expires, so the retry has to re-resolve the redirect rather than refetch a dead URL.[dispatcher.github_retries]). Widening what may be retried would make a duplicate side effect a setting.ddev/src/ddev/utils/github_async/AGENTS.mddocuments the layer boundary so the next change lands in the right one.Motivation
Closes AI-7115.
Dispatcher runs for hours and makes thousands of GitHub calls, and until now any failure that was not rate limiting failed on the first attempt. That gives a single blip more power than it should have.
TaskTestRunnerpollsget_workflow_runfor the whole life of a batch inside atry/finallywith noexcept, so one transient 500 aborts the batch, closes its check run as cancelled and throws away the results of every test in it. Other calls swallow the failure and quietly degrade instead: a failedlist_workflow_jobsreturns an empty job list, so job correlation silently loses data.Both get worse as we scale up: more batches and more polling mean more chances to hit the one blip that costs a whole batch of test results. Retrying is also a precondition for trusting the run report, since a report that is missing jobs because of a dropped connection is worse than one that is late.
No task behaviour changes here. Retries only make those paths less likely to fire, and a failure that outlives the ladder surfaces exactly as it does today.
Notes for review
ExecutionState.RETRYING,BatchProgress.retrying_jobs). This is unrelated and only concerns HTTP requests.test_no_retry_on_transport_errorbecametest_the_rate_limit_layer_does_not_retry_a_transport_errorand now calls_rate_limited_request. The property still holds for that layer, but at client level a GET transport error is now retried on purpose.GitHubAuthenticationError, which the guard refuses, so a real denial still fails immediately. Tested both ways.staminalogger, so retries are visible even with no logger injected. Turning it off is global and would also silence the unrelatedstamina.retryinddev/e2e/agent/docker.py, so I left it and documented it. Say the word if you want the client to be the only voice.Review checklist (to be filled by reviewers)
qa/requiredif this PR needs QA validation, orqa/skip-qaif it does not. Exactly one of the two is required.backport/<branch-name>label to the PR and it will automatically open a backport PR once this one is merged